Papers with multilingual learning
Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER (P18-2)
Copied to clipboard
| Challenge: | Existing approaches to improve NER performance add training data from one or more assisting languages to the primary language. |
| Approach: | They propose a metric based on symmetric KL divergence to filter out highly divergent training instances in the assisting language. |
| Outcome: | The proposed method improves NER performance in many languages, including those with limited training data. |
75 Languages, 1 Model: Parsing Universal Dependencies Universally (D19-1)
Copied to clipboard
| Challenge: | UDify is a multilingual multi-task model that can predict universal part-of-speech, morphological features, lemmas, and dependency trees. |
| Approach: | They evaluate UDify, a multilingual multi-task model capable of predicting universal part-of-speech, morphological features, lemmas, and dependency trees simultaneously for all 124 Universal Dependencies treebanks across 75 languages. |
| Outcome: | The proposed model can predict universal part-of-speech, morphological features, lemmas, and dependency trees for all 124 treebanks across 75 languages. |
X-SRL: A Parallel Cross-Lingual Semantic Role Labeling Dataset (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing multilingual SRL datasets contain disparate annotation styles or come from different domains, hampering generalization in multilingual learning. |
| Approach: | They propose to automatically construct an SRL corpus that is parallel in four languages with unified predicate and role annotations that are fully comparable across languages. |
| Outcome: | The proposed method improves performance for English SRL in weaker languages. |
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)
Copied to clipboard
| Challenge: | Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us . |
| Approach: | They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages. |
| Outcome: | The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning. |
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing multimodal corpora lack the ability to be used in multilingual or non-English scenarios. |
| Approach: | They extend a Flickr30k Entities image-caption dataset with Japanese translations to provide a multilingual corpus. |
| Outcome: | The proposed dataset is the first multilingual image-caption dataset with Japanese translations. |
Hyper-X: A Unified Hypernetwork for Multi-Task Multilingual Transfer (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing multilingual models cannot fully leverage training data when it is available in different task-language combinations. |
| Approach: | They propose a single hypernetwork that unifies multi-task and multilingual learning with efficient adaptation. |
| Outcome: | The proposed model achieves the best or competitive gain when a mixture of multiple resources is available while being significantly more efficient than existing models. |
Don’t Go Far Off: An Empirical Study on Neural Poetry Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | despite improvements in machine translation quality, automatic poetry translation remains a challenging problem . et al., a study of automatic poetry translators shows that multilingual fine-tuning on poetic data outperforms bilingual fine-timing on non-poetic text . |
| Approach: | They propose to use poetic parallel corpora for 6 languages to study poetry translation . they find that multilingual fine-tuning on poetic data outperforms bilingual fine-uning . |
| Outcome: | The proposed model outperforms bilingual and multilingual models on poetic data . the proposed model is based on a parallel dataset of poetry translations for several languages . |
Polyglot Prompt: Multilingual Multitask Prompt Training (2022.emnlp-main)
Copied to clipboard
| Challenge: | a monolithic framework for multilingual learning can be used without any task/language-specific module. |
| Approach: | They propose a framework to exploit prompting methods for learning a unified semantic space for different languages and tasks with multilingual prompt engineering. |
| Outcome: | The proposed framework can learn tasks from different languages in a monolithic framework without any task/language-specific module. |
Multilingual Transfer Learning for Children Automatic Speech Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in automatic speech recognition (ASR) systems have been criticized for high acoustic variability and limited amount of available training data. |
| Approach: | They propose a two-step training strategy that uses multilingual learning followed by language-specific transfer learning to generalize children's speech. |
| Outcome: | The proposed training strategy outperforms single language training and multilingual and transfer learning alone in English. |
Typology Guided Multilingual Position Representations: Case on Dependency Parsing (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent multilingual models benefit from strong unified semantic representation models, but conflicting linguistic regularities may break the effectiveness of word position features in multilingual learning. |
| Approach: | They propose to combine prior knowledge from typology features and existing position vectors to create a position generation network which combines prior knowledge of a language's position space and typological characterization. |
| Outcome: | The proposed model can achieve the best multilingual parsing results by combining prior knowledge from typology features and existing position vectors. |
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in natural language processing (NLP) have led to significant breakthroughs in the field. |
| Approach: | They evaluate ChatGPT over multiple tasks with diverse languages and large datasets to provide more comprehensive information for multilingual NLP applications. |
| Outcome: | The proposed model can process and generate texts for multiple languages due to its multilingual training data. |